Back

in silico Plants

Oxford University Press (OUP)

Preprints posted in the last 90 days, ranked by how well they match in silico Plants's content profile, based on 27 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
GreenSloth: a curated database and executable platform for mechanistic photosynthesis models

Corvest, E.; van Aalst, M.; Nies, T.; Nguyen, Q. H.; Ebeling, J.; Strauch, M.; Cisse, E.-H. M.; Hassan, T.; Matuszynska, A.

2026-07-26 systems biology 10.64898/2026.07.22.740007 medRxiv
Top 0.1%
19.3%
Show abstract

Mechanistic models of photosynthesis have expanded substantially over the past decades, covering processes from light reactions to carbon fixation. However, these models remain fragmented across the literature, inconsistently implemented, and difficult to reproduce or reuse, limiting their adoption beyond the research group that developed them. Here, we present GreenSloth, a freely accessible web-based database of 22 published mechanistic photosynthesis models, reimplemented as standardized, executable Python objects within MxlPy, an open-source framework for mechanistic biological modeling. Although the database is primarily designed for dynamic mechanistic models formulated as ordinary differential equations, the current implementation also includes the fields most widely cited steady-state mechanistic model and its variants. GreenSloth provides a structured environment for model discovery, comparison, and reuse, addressing reproducibility challenges in the field and enabling integration into emerging hybrid modeling approaches. It is also interactive: each model runs directly in the browser, with no installation, environment setup, or programming required. The resource is openly accessible and designed for long-term community maintenance, hoping to position itself as foundational infrastructure for the photosynthesis modeling community. Database URLhttps://greensloth.rwth-aachen.de/

2
Integrating carbon utilization and transport processes into a crop growth model enables the prediction of emergent soybean carbon allocation behavior

Piao, X.; Lochocki, E. B.; McGrath, J.; Matthews, M. L.

2026-08-28 plant biology 10.64898/2026.08.27.747615 medRxiv
Top 0.1%
18.8%
Show abstract

Accurately modeling carbon (C) allocation is essential for predicting crop yield and the performance of new cultivars in various environments. Most crop models allocate C empirically, using fixed partitioning tables or harvest indices that prescribe allocation without representing the underlying physiology, limiting their predictive power under novel conditions. A mechanistic alternative, in which C allocation emerges from local utilization and transport, could instead respond dynamically to environmental changes, source-sink perturbations, and organ-level trait modifications. To achieve this design, we integrated a utilization-transport-resistance (UTR) allocation model into the Soybean-BioCro crop growth modeling framework. We calibrated and validated the model using organ biomass data from two soybean cultivars grown at two CO2 levels over eight seasons, achieving accuracy comparable to partitioning-based models while predicting more reasonable carbon allocation fractions. Further, the UTR-BioCro model predicted leaf and stem total nonstructural carbohydrate concentrations with reasonable accuracy compared to experimental measurements across the 2022 growing season. A local sensitivity analysis of the model parameters indicated that the onset of reproductive growth influenced yield more strongly than utilization or transport parameters suggesting the timing of this transition as a potential target for crop improvement. Finally, the UTR-BioCro model reproduced yield responses to source-sink perturbations including shading and pod removal, and captured the qualitative response to defoliation without requiring scenario-specific tuning as most partitioning approaches require. By grounding C allocation in physiological mechanisms, this work provides a foundation for predicting crop responses across diverse environments and engineered traits, supporting crop improvement for a changing environment.

3
Evaluating crop models for future climate scenarios: wheat yield predictions using APSIM and STICS under combined CO2, warming, and water deficit conditions

Severini, A. D.; Gawinowski, M.; Bancal, M.-O.; Launay, M.; Deswarte, J.-C.; Chenu, K.

2026-06-10 plant biology 10.64898/2026.06.07.730737 medRxiv
Top 0.1%
18.4%
Show abstract

Crop models are essential for predicting climate change impacts on agriculture, yet their validation under multi-stress conditions remains limited. This study evaluated two widely-used wheat models, APSIM and STICS, using data from three Free-Air CO2 Enrichment (FACE) experiments (USA, Germany, Australia) combining elevated CO2 (eCO2), water deficit, and warming. Environmental characterisation using simulation-based stress indices revealed that intended "controls" frequently experienced hidden heat and water stress, meaning models were calibrated on crops already undergoing physiological adjustments. Evaluation of simulated yield and components revealed a clear hierarchy in prediction errors (RRMSE): unlimited conditions (3-9%) < single stress (4-27%, with a need to improve response to heat stress) < combined stress (17-123%). Elevated CO2 generally increased prediction uncertainty for crops experiencing water stress. Our results suggest that current stress functions from the models fail to capture the synergistic coupling between drought and heat stress. This highlights the urgent need for more mechanistic modelling to improve the reliability of climate change impact assessments.

4
Knowledge-guided Bayesian optimization using pre-trained LLMs speeds up the identification of superior genotypes from germplasm collection

Hamazaki, K.; Tsuda, K.

2026-07-02 bioinformatics 10.64898/2026.06.28.735149 medRxiv
Top 0.1%
15.2%
Show abstract

Background: Germplasm collections contain wide genetic diversity that is valuable for plant breeding, but conducting phenotypic evaluation for all genotypes in field trials is rarely feasible. Bayesian optimization offers a way to decide, season by season, which genotypes to cultivate in order to identify superior genotypes with fewer evaluations. However, standard Bayesian optimization commonly starts from randomly selected genotypes and mainly relies on surrogate models built from marker genotype information, while the text-based passport information that accompanies germplasm is not fully used. We examined whether pre-trained large language models can provide prior knowledge that improves these decisions in germplasm evaluation. Results: We constructed a large-language-model-guided Bayesian optimization framework that introduces large language models into two parts of the Bayesian optimization workflow. In zero-shot warmstarting, a large language model proposes initial genotypes using passport information such as cultivar name, country of origin, and subpopulation, optionally together with principal component scores derived from genome-wide single-nucleotide-polymorphism markers. In addition, we evaluated a large-language-model-based surrogate model that predicts phenotypic values for untested genotypes using in-context learning from previously evaluated genotypes. Using a rice germplasm panel and two target traits (seed number per panicle for maximization and protein content for minimization), we compared strategies. For seed number per panicle, zero-shot warmstarting with a general-purpose instruction-following model reduced the number of evaluated genotypes needed to reach the best genotype, whereas improvements were small for protein content. When genomic information was available, Gaussian-process-based Bayesian optimization was the strongest overall approach, while the large-language-model-based surrogate model outperformed random baselines and was competitive in some settings. When genomic information was not available, predictions based on passport information improved efficiency compared with fully random strategies. Conclusions: Pre-trained large language models can inject useful agronomic knowledge into Bayesian optimization for germplasm evaluation, particularly by improving early-stage genotype selection, and can also support optimization when genomic information is unavailable. As models better handle long genomic sequences together with passport information, large-language-model-guided Bayesian optimization may become a practical and explainable decision-support approach for agricultural optimization.

5
A general mathematical framework for modelling subnetworks of the nuclear auxin pathway

Shuttleworth, J. G.; Chan, E.; Welch, T.; Bhosale, R. G.; Bishopp, A.; Farcot, E.

2026-08-07 plant biology 10.64898/2026.08.06.742982 medRxiv
Top 0.1%
10.7%
Show abstract

Auxins are a family of plant hormones involved in various processes across plant tissues and species. The Nuclear Auxin Pathway (NAP) consists of interacting transcription factors (ARFs) and repressors (Aux/IAAs), which govern an individual cells response to changes in auxin concentration. These components are present in all land plants, and many species possess multiple copies of each signalling component. We present a general framework for ODE-based models of NAP submodules with the flexibility to model the promotion and repression of target genes by any combination of transcriptional regulators. We analyse published data and show that auxin treatment in Arabidopsis thaliana roots triggers a range of characteristically distinct temporal response profiles--for both target genes and the signalling components themselves. Using our modelling framework, we recapitulate aspects of this behaviour by presenting examples of real and theoretical NAP subnetworks, and by analysing the effect that these network dynamics have on auxin-mediated transcriptional responses. This work demonstrates the utility of our modelling framework as a general-purpose tool for understanding the function of certain protein-protein and protein-DNA interactions through their effects on the NAP. This exploration of the rich dynamics of more complex signalling pathways promises to advance our understanding of the NAP.

6
A guaranteed-convergence algorithm for coupled leaf photosynthesis–transpiration–stomatal conductance models

Masutomi, Y.;Kobayashi, K.

2026-07-08 Plant Biology 10.64898/2026.06.24.734164 medRxiv
Top 0.1%
10.2%
Show abstract

The photosynthesis-transpiration-stomatal conductance (An-E-gs) model framework is widely used for estimating photosynthesis, transpiration, and stomatal conductance in plants. The model equations are solved by numerical iteration, and the converged model values are deemed the solution. However, there has been no general guarantee that the iterative procedure converges to a solution or that the procedure leads to convergence. Building on the recent proof of the existence of a unique set of solutions, we herewith propose a numerical algorithm that is guaranteed to converge to the solution for the An-E-gs model framework. We first analytically prove that the proposed algorithm necessarily converges to a solution. We then demonstrate the convergence across contrasting combinations of leaf temperature, relative humidity, light, atmospheric CO2, and wind speed. We further demonstrate rapid convergence with the algorithm: no more than ca. 10 iterations for approximately 10-3 mol CO2 m-2 s-1 precision in net photosynthesis and no more than ca. 20 iterations for 10-7 mol CO2 m-2 s-1 precision. By guaranteeing convergence to the solution, this algorithm eliminates concerns about nonconvergence in leaf gas-exchange calculations and is expected to serve as a robust foundation for a range of studies from leaf-level gas exchange to global-scale carbon and water cycle dynamics.

7
From Diverse Prior Knowledge to Mechanistic Causal Network Using PSoup: A Case Study in Shoot Branching

Mitsanis, C.; Fortuna, N. Z.; Beveridge, C.

2026-08-10 plant biology 10.64898/2026.08.07.743620 medRxiv
Top 0.1%
7.7%
Show abstract

Mechanistic models of plant regulatory networks typically require extensive parameterization, limiting their generalisation and scalability. Here we present a parameter-free, topology-driven model of shoot branching that predicts phenotypic outcomes from network structure alone. We constructed a signed, directed causal network by distilling regulatory relationships from the published literature spanning many laboratories, species, years, data types, and methodological frameworks. This extracted the essential logic of the system, consistent with developmental-biological reasoning and anchored in empirical evidence. Using PSoup, which automatically translates network topology into algebraic equations, the model propagates information across the network and predicts the qualitative direction of change relative to a defined baseline, mirroring the comparative framework of biological experiments. The pipeline, from network construction through automated equation generation to prediction, is transparent and reproducible. Trained against branching phenotype data with 78 diverse perturbations spanning genetic mutations and hormone treatments, the model achieved 86% accuracy in predicting branching direction. On an independent test set of 84 perturbations measuring bud release and gene expression at nodes not used during training, accuracy reached 75%. The approach highlighted deficiencies in our understanding of the topology of the network around SMXL 6/7/8 and ABA nodes. Other errors came mainly from modelling choices, such as the threshold for scoring a node as changed relative to baseline. Beyond shoot branching, this work demonstrates a general strategy for synthesizing biological knowledge into validated predictive networks, providing a foundation for both applied breeding and the advancement of fundamental biology.

8
Gravitropism Shapes the Pareto Front of Root System Architecture

Garza, A.; Altman, K.; Koenig, A.; Richards, K.; Rahmati-Ishka, M.; Lobet, G.; Julkowska, M. M.; Chandrasekhar, A.

2026-07-17 plant biology 10.64898/2026.07.16.738932 medRxiv
Top 0.1%
6.3%
Show abstract

The root systems of wild tomatoes (S. Pimpinellifolium) can be understood as biological networks in which the lateral roots branch from a single main root and together balance two competing objectives: minimizing the material cost of building the network (wiring cost) and minimizing the transport time from the root tips to the shoot (conduction delay). Our prior work showed that S. Pimpinellifolium root architectures cluster near the Pareto-optimal front defined by these two objectives, with morphological diversity resolving into four qualitative topologies ((Chandrasekhar and Julkowska, 2022)). That framework assumed lateral roots grow as straight lines - ignoring gradual onset of lateral root gravitropism, the tendency of roots to curve toward the gravity vector. Because curved trajectories are longer than straight lines, gravitropism directly increases both wiring cost and conduction delay, constraining which architectures are physically realizable and thereby reshaping the Pareto front itself. Here we extend the model to explicitly incorporate lateral root gravitropism, producing predicted architectures that align much more closely with observed S. Pimpinellifolium root systems. We present a computational method to infer gravitropic sensitivity directly from anatomical tracing data - without reorientation assays - and apply it to 2423 arbors across different root topologies, growth conditions, and hormone treatments. Incorporating gravitropism reveals variation invisible to the straight-line model:notably, lateral roots show reduced gravitropic sensitivity under salt stress, mirroring a phenomenon previously described only for main roots and overlooked in lateral roots until now.

9
Dynamic balance of sparse flux vectors for efficient simulation of culture dynamics and metabolic network reduction

Tapia García, I.; Torrealba, C.; Luna, R.; Pérez-Correa, J. R.; Saa, P. A.

2026-06-22 bioinformatics 10.64898/2026.06.17.733012 medRxiv
Top 0.1%
5.7%
Show abstract

Dynamic Flux Balance Analysis (DFBA) enables simulation of microbial culture dynamics under changing environmental conditions, but remains computationally expensive for tasks such as parameter calibration and fermentation optimization when applied using genome-scale metabolic models (GEMs). To address this challenge, we introduce Dynamic Flux Vector Balancing (DFVB), a reformulation of DFBA that solves an equivalent problem using a pre-computed, sparse basis of flux solutions that reduces the dimensionality of the internal optimization problem without information loss. Notably, DFVB provides a compact, interpretable representation of flux states that can readily identify dynamically inactive pathways and enable simulation-based automatic metabolic network reduction. We showed that DFVB produces the same culture dynamics as DFBA across multiple model scales and conditions, and identifies inactive reactions more accurately than Flux Variability Analysis (FVA) when compared to transcriptomic data profiles. Furthermore, computational performance analyses demonstrated that integrating DFVB with solver warm-start strategies and model reduction enhances computational efficiency relative to DFBA, yielding up to 3-fold reductions in simulation time for large-scale metabolic models. Finally, kinetic parameter estimation of culture dynamics with DFVB in two fermentation scenarios using a large-scale yeast GEM reached equal or higher prediction fidelity and narrower confidence intervals than DFBA, indicating improved parameter identifiability and robustness. Together, these results position DFVB as a scalable, robust, and biologically coherent framework for dynamic metabolic modeling, easing the integration of GEMs for culture dynamics simulation.

10
TWIST: A diagnostic framework for representing tree water deficit dynamics in process-based forest models

Ziegler, Y.;Labenski, P.;Thurner, M.;Krejza, J.;Sigut, L.;Ruehr, N.;Grote, R.

2026-06-16 Plant Biology 10.64898/2026.06.15.732331 medRxiv
Top 0.1%
5.5%
Show abstract

Dendrometer-derived tree water deficit (TWD) contains physiologically rich information and is increasingly used to monitor tree drought stress, yet process-based forest models rarely include a directly comparable representation of TWD dynamics. Existing hydraulic models can represent internal water storage in detail, but their parameter demands limit broader application. Here, we introduce the Tree Water Imbalance and Storage Tracker (TWIST), a parsimonious and physiologically interpretable framework that derives volume-based TWD dynamics. The module is driven by transpiration and relative soil water content and uses three empirical parameters to control transpiration-driven internal water depletion, deficit refilling, and additional soil-water uptake limitation. It also derives relative tree water content (RWCtree) from the simulated deficit and an estimate of the available internal water pool. We tested TWIST by coupling it to the process-based ecosystem model LandscapeDNDC. Parameters were optimized for 2018 and evaluated independently for 2019-2024 against normalized dendrometer-derived TWD at a Czech beech site. Simulated TWD trajectories broadly agreed with observed daily and seasonal dynamics, while RWCtree translated them into a physiologically interpretable proxy for internal dehydration. TWIST demonstrated capability to reproduce key TWD drought-response patterns, including diurnal depletion-replenishment cycles, reduced nocturnal rehydration with declining soil moisture, and progressive deficit accumulation. By representing TWD and RWCtree as diagnostic model outputs, TWIST makes dendrometer-derived drought-stress information more directly usable in forest models. It thereby provides a practical basis for linking tree-level drought-stress signals with stand-level simulations and, potentially, remotely sensed indicators of canopy water status.

11
Beyond climatic drought indices : an hydraulic approach to quantifying forest water stress

Cochard, H.

2026-07-15 plant biology 10.64898/2026.07.13.738371 medRxiv
Top 0.1%
5.4%
Show abstract

The article introduces a new Forest Stress Index (ISF) based on a plant hydraulic modelling approach rather than classical climatic drought indices. Unlike other index like scPDSI or SPEI, ISF is grounded in xylem embolism dynamics simulated with the mechanistic SurEau model. The goal is to better link climatic anomalies to tree physiological functioning and mortality risk. ISF is defined using a locally adapted ideotype characterized by an optimal P50 value under a reference hydraulic functioning threshold. Simulations are performed across Europe and France using multiple climate datasets. The index is robust to model parameterization choices and assumptions about plant functional traits. Results show strong spatial and temporal consistency and significant correlations with SPEI and scPDSI. However, ISF more strongly highlights extreme drought years and exhibits a more skewed distribution. Future projections under SSP5-8.5 indicate a widespread increase in hydraulic stress with strong regional contrasts. Overall, ISF provides a mechanistic and complementary drought indicator more directly linked to forest mortality processes.

12
From Field Photosynthesis to Genetic Architecture: Insights from the First Dedicated Photosynthesis Hackathon

Matuszynska, A.; Sansa, O.; Adekoya, F. J.; Akinyemi, O. O.; Anokye, E.; Bashir, O. B.; Boyny, Z. Z. F.; Chukwuka, M. K.; Corvest, E.; Dada, A. O.; DellAcqua, M.; Ehemba, G. L.; Finkbeiner, A. J.; Hamabwe, S.; Hodehou, D. A. T.; Kacheyo, O.; Kamfwa, K.; Mhango, K. J.; Abdullahi, W. M.; Munduwe, G.; Ntukidem, S.; Obisesan, O. K.; Odesina, I. S.; Ogechi, N.-U.; Olaoye, O. D.; Olayinka, M. M.; Osei-Bonsu, I.; Rilwan, K. O.; Stival, L.; Tehar, Z.; Tende, R. M.; To, J.; Ugochukwu, U. K.; Unger, A.; van Aalst, M.; Vrbic, D.; Zhang, C.; Theeuwen, T. P. J. M.; Kramer, D. M.; Kromdijk, J.

2026-08-17 plant biology 10.64898/2026.07.24.740625 medRxiv
Top 0.1%
5.3%
Show abstract

Photosynthesis is among the most consequential yet genetically complex traits in crop plants, and translating its natural variation into actionable genomic targets remains a central challenge for breeding climate-resilient varieties. To start addressing this, researchers are generating increasingly large, multi-environment field photosynthesis datasets. Yet, these data have been structurally under-analysed since their inception. Here we report the outcomes of the first dedicated hackathon focused on computational mining of such field data held in Accra, Ghana, in March 2026. Bringing together data scientists, plant physiologists, geneticists, and breeders from Europe and Africa, these interdisciplinary teams used photosynthetic data collected with hand-held fluorometers to genome-wide marker data across four crop species: cowpea (Vigna unguiculata), barley (Hordeum vulgare), common bean (Phaseolus vulgaris), and potato (Solanum tuberosum). Despite using different species and methods, independent teams identified the same three key findings. First, mechanism-informed feature engineering and dynamic modelling recover genetic signals that are not detected or discarded in standard analysis pipelines, resulting in traits with improved heritability and meaningful associations with yield. Secondly, machine learning methods proved effective at uncovering genetic associations, with temporally resolved features substantially outperforming single time-point measurements. Third, raw chlorophyll fluorescence and absorbance traces consistently contained more information and predictive power than the extracted parameters currently used. A defining feature of this event was having experimentalists and data scientists working together, enabling AI approaches to be grounded in domain knowledge and biological mechanisms rather than relying on data alone.

13
Identification of environmental factors and growth stages in the prediction of fibre yield and fibre quality traits in rain-grown cotton

Feng, Q.; Rafter, P.; Wilson, I.; Li, Z.; Conaty, W.

2026-06-18 bioinformatics 10.64898/2026.06.14.732217 medRxiv
Top 0.1%
5.3%
Show abstract

ContextUnderstanding how and when environmental conditions influence overall crop performance is crucial for optimising the development of genotypes to a specific breeding target environment. We focused on economically important traits of Australian rain-grown cotton including fibre yield and quality traits, which have not been investigated comprehensively. The aim of the study was to identify relevant environmental factors, and the timing and extent of their impact on rain-grown cotton production. MethodsWe used a data driven approach to analyse the relationship between ten climate related environmental factors across various plant growth stages and eight fibre yield and quality traits, using a large-scale field dataset of 9,283 records collected over 23 years at 4 locations, with 53 unique year-location combinations. We applied eight complementary statistical models including stepwise, penalised and Bayesian linear regression, regression-tree based ensemble methods and deep learning frameworks to (1) select the most essential environmental covariates affecting rain-grown cotton production, and (2) evaluate the predictive performance of these models. ResultsThe environmental impacts on rain-grown cotton production were trait and growth-stage specific. Number of rainy days and solar radiation were identified as the most influential environmental factors for fibre yield traits, vapour pressure deficit at maximum daily temperature was the most influential factor for majority of fibre quality traits. However, each analysed trait was influenced by multiple environmental factors across multiple growth stages (rather than a single factor or a single growth stage). These influential covariates explained a wide range of variation in the traits, accounting for 5.8% to 68.2%. Using the best-fit random forest model, our findings revealed non-linear relationships between key environmental covariates and the traits. ConclusionsEnvironmental factors at different rain-grown cotton growth stages are key determinants for the performance of end-of-season fibre yield and fibre quality parameters. These findings highlight the need to account for environment conditions when developing cotton varieties optimised for rain-grown production systems. Potential strategies are proposed whereby these key environmental factors can be used to increase the rate of genetic gain in rain-grown cotton production systems. ImplicationsThe results of this study will be crucial for future genetic evaluations and analyses of genotype-by-environment interaction effects in rain-grown cotton, which must account for the influence of the environment on plant performance. Furthermore, these methods can be applied to other species to identify critical growth stages and environmental factors which most influence crop performance.

14
Root shape phenotyping using three-dimensional root system vector data in rice

Teramoto, S.;Uga, Y.

2026-06-15 Plant Biology 10.64898/2026.06.15.732232 medRxiv
Top 0.1%
3.3%
Show abstract

PurposeAlternation of root distribution in the soil is a method for improving root system architecture (RSA) in crops. If the root is straight, root distribution has been altered by regulating the angle of root growth in the vertical direction. However, the root shape should be curved and winding. This study aimed to define parameters reflecting the actual root shape. MethodsWe used three-dimensional vector data of rice (Oryza sativa L.) RSA derived from an X-ray computed tomography image to compile two sets of two-dimensional vector data for horizontal and vertical components. In the vertical component, we defined the dropping angle{theta} d, which is calculated assuming that the rice roots are bent upward. In the horizontal component, we defined the polar angle{theta}{rho} , which is the direction in which the roots grow when viewing the plant from above, and the winding degree log {sigma}w, which is calculated by assuming that root elongation is in a random walk. ResultsAssuming that{theta}{rho} is distributed uniformly in all varieties, there should be no varietal differences in this angle. We measured{theta} d and log {sigma}w of three rice varieties with different root distributions: shallow, deep, and intermediate. We found significant varietal difference in{theta} d and log {sigma}w. ConclusionsWe have shown that{theta} d and log {sigma}w are useful parameters for comparing RSA in rice varieties. We have named this methodology RSAparam3D, and it is freely available to researchers.

15
Root phenotypic plasticity improves yield stability when directed toward an adaptive integrated phenotype

Lopez-Valdivia, I.; Tawale, A. B.; Schierenbeck, M.; Sandoni, D.; Jones, D. H.; Kirschner, G. K.; Schneider, H. M.

2026-08-11 plant biology 10.64898/2026.08.10.744026 medRxiv
Top 0.1%
3.2%
Show abstract

Root phenotypic plasticity is often proposed to improve crop performance under stress, yet it remains unclear how much plasticity is beneficial and whether adaptive responses require changes across many traits or adjustments in few specific traits. Using public data of 6,500 field-grown maize and barley plants, this study examined the extent and distribution of root plasticity, and when it is associated with yield stability. We quantified root plasticity across nine anatomical and architectural traits using complementary statistical models and applied a feature-discovery framework to identify the drought-associated optimal integrated phenotypes and determine whether plasticity toward these phenotypes improved yield stability. More plasticity did not mean greater yield stability. Neither the number of plastic traits nor the magnitude of plastic responses predicted yield stability. Rather, we identified species-specific high-yielding, stable integrated phenotypes defined by distinct trait configurations. Critically, genotypes whose plastic responses moved their root phenotype toward these targets achieved greater yield stability, whereas movement away from them was associated with lower stability. Root plasticity is adaptive when it shifts root phenotypes towards an optimal integrated phenotype. These findings show that the value of plasticity depends on the trajectory of phenotypic change rather than its magnitude alone.

16
Additive Effects Dominate Legume Responses to Combined Heat and Drought Stress: A Quantitative Review

Meijer, L.; Chenu, K.; Smith, M. R.; Van Haeften, S. R.; Sadras, V.

2026-08-13 plant biology 10.64898/2026.08.12.744551 medRxiv
Top 0.1%
3.2%
Show abstract

Concurrent exposure to heat and drought stress compromises legume productivity, yet their combined effects are rarely quantified systematically. We compiled a database of 18 studies covering seven legume species. From these, we extracted 929 physiological, biochemical, and yield-related traits and calculated actual-to-additive ratios to classify heat-drought interactions as antagonistic (ratio < 1), additive (ratio = 1), or synergistic (ratio > 1). Additive heat-drought relationships accounted for 59 % of all classifiable observations, 37% relationships were antagonistic, and 4% synergistic. The relationship varied with species, genotype, trait, and experimental conditions highlighting the complexity of combined abiotic stress effects. The results challenge the common assumption that concurrent stresses invariably exacerbate damage and underscore the need for more realistic, quantitatively defined stress treatments as well as frameworks that integrate trait-level responses into predictive models of crop growth and development. Our synthesis provides a quantitative foundation to understand legume phenotypes under the increasingly frequent co-occurrence of heat and drought stress and identifies research areas where further work is needed to improve insight into combined stress responses. HighlightsO_LICombined heat and drought responses were mainly additive or antagonistic. C_LIO_LIEvidence is biased toward few legumes and controlled environments. C_LIO_LIField-based, multi-species studies are needed to identify adaptive traits. C_LI

17
Plant DNA Designer: A Computational Framework for Multi-Objective Codon Optimisation and Synthetic Gene Design in Crop Biotechnology

k, D.; H, S.

2026-08-06 bioinformatics 10.64898/2026.08.02.742271 medRxiv
Top 0.1%
2.4%
Show abstract

Synthetic gene design for plant transformation requires simultaneous optimisation of multiple, often competing, molecular objectives: translational efficiency, mRNA structural accessibility, codon-pair compatibility, regulatory safety, and species-specific expression context. Existing tools address these objectives in isolation, typically maximising a single metric such as the Codon Adaptation Index (CAI) and neglecting the broader determinants of in-plant expression. We present Plant DNA Designer (PDD), a web-based platform that integrates a 19-objective genetic algorithm with expression-cassette co-design, clade-aware translation-initiation logic, ribosome-velocity trajectory shaping, CRISPR guide-RNA design, and multi-gene pathway balancing across 18 crop species spanning monocot and dicot clades -- each using its own measured codon-usage table from the Kazusa Codon Usage Database. We benchmark PDD against faithful reproductions of the published algorithms of five external tools (JCat/OPTIMIZER/ATGme, IDT, TISIGNER, a CAI+GC heuristic, and a random floor) across six validated rice effector proteins. PDD is the only strategy that holds every objective within acceptable bounds at once: it reduces transgene safety liabilities from 2.3-3.5 to 0.0, and cuts deviation from a 50 % GC synthesis target from 21.8 to 4.0 percentage points, while raising codon harmony from 0.42 to 0.77 -- at a deliberate, moderate cost in raw CAI (0.79 vs 1.00). Consistent with a fair comparison rather than a strawman, a dedicated single-objective tool (IDT) still outperforms PDD on its own axis (harmony 0.93). We anchor the two central proxies against real biology: on 456 real rice genes, CAI and the wobble-weighted tAI are significantly higher in highly-expressed ribosomal-protein genes than in the genomic background (Mann-Whitney p [&le;] 10-; tAI AUC 0.75) and correlate at Spearman{rho} = 0.93. Beyond this expression-class anchor, the reported design metrics are in-silico proxies, not wet-lab yield measurements. PDD is released as open-source software under an MIT licence and is freely accessible as a FastAPI web application.

18
Impact of Reduced Chlorophyll Levels in Leaves on Soybean Yield, Seed Composition, Pod/Seed Photosynthesis, and Chlorophyll Levels in Pod and Seed Tissues

Jones, S. I.; Stutz, S. S.; Atalay, E.; Wang, Y.; Ort, D. R.; Cho, Y. B.

2026-08-19 plant biology 10.64898/2026.08.14.744892 medRxiv
Top 0.1%
2.3%
Show abstract

Soybean, a widely cultivated leguminous crop valued for its protein, amino acids, and oil, faces the challenge of maintaining protein levels, which have an inverse correlation with yield. Reducing leaf chlorophyll levels could increase seed protein levels without compromising yield; however, this is yet to be tested. Therefore, to understand the impacts of low chlorophyll mutations on soybean yield and seed composition, we screened and compared 25 low chlorophyll soybean mutants to their 11 dark green parents. PI548210 (Lincoln mutant) demonstrates a higher concentration of protein without affecting yield compared to its dark green parent PI548362 (Lincoln), suggesting it as a good candidate for further large-scale field trials. PI547555 (Y11/y11, Clark mutant) demonstrates a lower concentration of oil without impacting yield, alongside lower gross photosynthesis, but with chlorophyll levels in the pod and seed tissues that are comparable to its dark green parent PI548533 (Clark). These findings are consistent with the oil concentration of the soybean being influenced by pod and seed photosynthesis, which is correlated with pod height and row spacing. Chlorophyll levels in the leaf do not necessarily correlate with those in the pod and seed of low chlorophyll mutants, possibly due to substantially lower expression of chlorophyll synthesis genes in the pod and seed. SIGNIFICANCEO_LIPI548210 (Lincoln mutant), one of twenty-five low chlorophyll soybean mutants, demonstrates a higher concentration of soybean protein without affecting yield compared to its dark green parent (Figure 1 and Table 1). C_LIO_LIPI547555 (Y11/y11, Clark mutant), a low chlorophyll soybean mutant, demonstrates a reduced concentration of soybean oil without impacting yield, alongside lower gross photosynthesis in pod and seed tissues compared to its dark green parent (Figures 3 and Table 2). These findings suggest that the oil concentration of the soybean is influenced by pod and seed photosynthesis, which is in turn influenced by pod height and row spacing (Figure 2). C_LIO_LIChlorophyll levels in the leaf do not necessarily correlate with those in the pod and seed of low chlorophyll mutants, possibly due to substantially lower expression of chlorophyll synthesis genes in the pod and seed (Figure 5-6). C_LI O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=84 SRC="FIGDIR/small/744892v1_fig1.gif" ALT="Figure 1"> View larger version (55K): org.highwire.dtl.DTLVardef@4282dcorg.highwire.dtl.DTLVardef@9d565forg.highwire.dtl.DTLVardef@1918292org.highwire.dtl.DTLVardef@1359b1_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1.C_FLOATNO Two low chlorophyll mutants are as healthy as their dark green parents. Lincoln and its low chlorophyll mutant, left; Clark and its low chlorophyll mutant, known as Y11/y11, right. It can be seen by eye that the plants have low chlorophyll (light green/yellow leaves) but a similar growth habit to their dark green parents. See Supplemental Figures 1-4 for contrast, where low chlorophyll mutants are stunted in growth compared to their dark green parents. C_FIG O_TBL View this table: org.highwire.dtl.DTLVardef@657ec9org.highwire.dtl.DTLVardef@166e75borg.highwire.dtl.DTLVardef@df23c7org.highwire.dtl.DTLVardef@1a60124org.highwire.dtl.DTLVardef@194ed96_HPS_FORMAT_FIGEXP M_TBL O_FLOATNOTable 1.C_FLOATNO O_TABLECAPTIONComparison of seed yield, weight, seed composition between low chlorophyll mutants and their dark green parents. ANOVA is used with linear mixed model (random effect = block, fixed effect = variety). Least squares mean is used to compare. For yield and seed composition, N=4 blocks. For leaf chlorophyll (SPAD), N=40. Yield is average yield per plant (g). n.s. = not significant. C_TABLECAPTION C_TBL O_FIG O_LINKSMALLFIG WIDTH=179 HEIGHT=200 SRC="FIGDIR/small/744892v1_fig3.gif" ALT="Figure 3"> View larger version (26K): org.highwire.dtl.DTLVardef@7a368aorg.highwire.dtl.DTLVardef@192b8f0org.highwire.dtl.DTLVardef@1abb738org.highwire.dtl.DTLVardef@89e978_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 3.C_FLOATNO Light response curve of low chlorophyll mutant (Y11/y11, PI547555) and its parent (Clark, PI548533). Rates of net and gross photosynthesis of low chlorophyll (white) and dark green parents (black) pods under field conditions. Each dot represents a value (n=4) {+/-}SE. We assumed that the seeds greatly inhibited the transmittance of light through the pod and used photosynthetic photon flux density for a single-side. C_FIG O_TBL View this table: org.highwire.dtl.DTLVardef@3f0528org.highwire.dtl.DTLVardef@16ba712org.highwire.dtl.DTLVardef@a5ab2aorg.highwire.dtl.DTLVardef@889254org.highwire.dtl.DTLVardef@3efa4f_HPS_FORMAT_FIGEXP M_TBL O_FLOATNOTable 2.C_FLOATNO O_TABLECAPTIONPod photosynthetic parameters for low chlorophyll mutant (Y11/y11, PI547555) and its parent (Clark, PI548533). Photosynthesis was measured 1 September through 15 September 2021 at the University of Illinois Energy Farm in Urbana, IL, USA. The statistical analysis was done using ANOVA with linear mixed model (alpha=0.05). N=4 {+/-} SEM for Clark and N=3 {+/-} SEM for Y11. C_TABLECAPTION C_TBL O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=130 SRC="FIGDIR/small/744892v1_fig2.gif" ALT="Figure 2"> View larger version (23K): org.highwire.dtl.DTLVardef@a36c26org.highwire.dtl.DTLVardef@1116c8forg.highwire.dtl.DTLVardef@ee5e61org.highwire.dtl.DTLVardef@1766712_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 2.C_FLOATNO Low chlorophyll mutant (Y11/y11, PI547555) and its parent (Clark, PI548533) differ in concentration of seed oil, which interacts with height of pod and row spacing. The box plots show the median (central line), the lower and upper quartiles (box) and the minimum and maximum values (whiskers). The statistical analysis was done using ANOVA with linear mixed model (n=3 blocks, alpha=0.05). Least squares mean is used to compare. N.s., non- significant in the analysis. A. Concentration of oil in low chlorophyll mutant seeds from the upper canopy decreased by 4% compared to the dark green parent (18.2% vs 19%) while there was no difference between them in the seeds from the lower canopy (20.2% vs 20.6%). B. Schematic layout of 2013 field setting showing two different row spacings. C. Concentration of oil in low chlorophyll mutant decreased by 2% in 38cm spacing (21.4% vs 22%) while there was no difference in 19cm spacing (21.3% vs 21.7%) in 2013 field. C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=162 SRC="FIGDIR/small/744892v1_fig5.gif" ALT="Figure 5"> View larger version (22K): org.highwire.dtl.DTLVardef@68e508org.highwire.dtl.DTLVardef@94a6ccorg.highwire.dtl.DTLVardef@152a187org.highwire.dtl.DTLVardef@1eae137_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 5C_FLOATNO (greenhouse). Correlation between the level of leaf chlorophyll (x-axis: SPAD reading) and the level of immature pod or seed chlorophyll (y-axis, mg/g DW). Line represents the linear regression model. R-squared is a coefficient of determination, the percentage of the response variable variation that is explained by the linear model. Pod is labeled by the fresh weight of seeds it contained. A. Level of chlorophyll of 25-100mg pod (n=18). B. Level of chlorophyll of 100-200mg pod (n=17) . C. Level of chlorophyll of 25-100mg seed (n=17). D. Level of chlorophyll of 100-200mg seed (n=20). C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=180 SRC="FIGDIR/small/744892v1_fig6.gif" ALT="Figure 6"> View larger version (28K): org.highwire.dtl.DTLVardef@167fd88org.highwire.dtl.DTLVardef@361472org.highwire.dtl.DTLVardef@786325org.highwire.dtl.DTLVardef@1b53855_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 6.C_FLOATNO Levels of gene expression in chlorophyll synthesis pathway. A. CHL common pathway genes; Glutamyl-tRNA reductase (GluTR). Glutamate 1- semialdehyde aminotransferase (GSA-AT). ALA dehydratase (ALAD). Uroporphyrinogen III synthase (UROS). Uroporphyrinogen III decarboxylase (UROD). Protoporphyrinogen IX oxidase (PPO). B. Mg branch; Mg-chelatase (Mgch). Magnesium-protoporphyrin IX monomethyl ester cyclase (MPEC). Protochlorophyllide reductase (POR). 3,8-divinyl protochlorophyllide a 8-vinyl-reductase (4VCR). Heme pathway; Ferrochelatase (FECH). Heme oxygenase (HO). Phytochromobilin synthase (HY). Data come from Severin et al (2010). RPKM, reads per kilobase per million mapped reads. DAF, days after flowering. The source seed is experimental line A81-356022 which was generated by introgressing G. soja (PI468916) into G. max (A81-356022). C_FIG

19
Enhancing predictive accuracy of yield traits in cassava through multi-trait genomic prediction

de Freitas, G. M.; Certuche, D. S.; Jannink, J.-L.; de Oliveira, E. J.; Garcia, A. A. F.

2026-07-06 genetics 10.64898/2026.07.01.735838 medRxiv
Top 0.1%
2.1%
Show abstract

Multi-trait genomic prediction offers a practical route to improve selection for costly, complex traits in clonally propagated crops such as cassava. In a Brazilian breeding panel of 1,078 cassava clones genotyped with 25,923 SNPs and phenotyped for six agronomic traits, we compared single-trait (ST) and multi-trait (MT) GBLUP models. Stage-wise mixed models produced BLUEs that fed into ST and MT-GBLUP. We tested five cross-validation schemes that mimic breeder realities: ST baseline (CV1); naive all-traits MT prediction for unphenotyped candidates (CV2); MT prediction using auxiliary trait phenotypes in the test set (CV3); and two sparse-phenotyping regimes with missingness by trait (CV4) or by clone (CV5) at 25%, 50%, and 75% levels. The main results were that, under the ST baseline (CV1), predictive ability ranged from 0.50 for DMC and 0.45 for FRY down to 0.13 for Le.Dis. A naive full MT model (CV2) performed approximately on par with ST-GBLUP. In contrast, MT designs (CV3) that included informative auxiliary traits, such as shoot yield and combinations with plant vigor and leaf disease severity, yielded small gains for DMC with predictive ability of approximately 0.51 (+2%), while FRY predictive ability increased to approximately 0.65 (+44%), accompanied by RMSE reductions for FRY up to approximately 13.5% (e.g. RMSE approximately 6.2). Sparse-phenotyping simulations (CV4/CV5) demonstrated that MT models sustain or even improve predictive ability under realistic missing-data regimes (PA {approx} 0.62 - 0.65). Selection concordance between MT and ST top-10% sets was generally high (>0.80), and MT configurations produced measurable improvements in expected selection response and genetic gain per cycle for several target traits. These results indicate that strategically implemented MT-GBLUP, using a small set of biologically and operationally informative auxiliary traits and optimized sparse phenotyping, can materially increase predictive accuracy and selection efciency for economically critical cassava traits while reducing phenotyping burden.

20
DeepPheno: A Deep Learning Framework for Linking Hyperspectral Imaging and SNP Genotypes in Lettuce

Okyere, F. G. G.; Mehrem, S. L.; Snoek, B. L.; Van den Ackerveken, G.; Abeln, S.

2026-07-10 plant biology 10.64898/2026.07.09.737449 medRxiv
Top 0.1%
2.1%
Show abstract

While whole genome sequencing captures millions of single nucleotide polymorphisms (SNPs) and hyperspectral imaging (HSI) enables non destructive plant phenotyping, integrating these modalities to link genotype to phenotype remains challenging due to their high dimensionality and non linearity. This study presents DeepPheno a deep learning framework that predicts SNP genotypes from HSI data, using model predictability as a proxy for genotype phenotype association. HSI data were acquired from 194 lettuce genotypes under field conditions. HSI data patches (20 x 20 pixels x 224 spectral bands) were used to train a hybrid CNN to predict the variant of a specific SNP. The framework was validated on SNPs with known phenotypic effects (anthocyanin, leaf serration, pale pigmentation), achieving high predictive performance (AUC ranging from 0.806 to 0.935), whereas models trained on randomly shuffled labels performed at chance (mean AUC {approx} 0.51). Extending the workflow to 50 randomly selected putatively neutral SNPs, most yielded low predictability, but two showed high performance (AUC > 0.76), suggesting uncharacterized genotype phenotype links. Explainable AI, including SHAP and Grad CAM, identified relevant spectral and spatial features driving these predictions, particularly the green and red edge wavelengths associated with pigment dynamics and leaf structure. These results establish a framework for understanding complex genotype phenotype interactions in plants and extracting these links from HSI data without predefining the exact trait values. It provides an avenue for high throughput trait discovery and description and extends the integration of image based phenomics with plant genetics.